Speed up permute propagation cleanup (#22161) - #22161
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/22161
Note: Links to docs will display an error until the docs builds have been completed. ✅ You can merge normally! (1 Unrelated Failure)As of commit 86beae5 with merge base d58fc25 ( FLAKY - The following job failed but was likely due to flakiness present on trunk:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
|
@apullin has exported this pull request. If you are a Meta employee, you can view the originating Diff in D114224762. |
This PR needs a
|
Summary: Speed up permute propagation cleanup. ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations: - Canonicalize view/permute chain collection now uses deque + membership set instead of list pop(0) and remove(), avoiding quadratic bookkeeping. - FuseDuplicateUsersPass deduplicates pending producer revisits while preserving same revisit behavior after fusions. - PropagateViewCopyPermutePass still retraces after each moved transform for metadata safety, but defers horizontal/vertical cleanup until full scan finds no more direct propagation moves. Preserves fixed-point behavior while avoiding expensive cleanup after every single moved transform. Behavior-preserving speedup, all existing pass tests pass. Differential Revision: D114224762
858c82b to
bb32005
Compare
Summary: Speed up permute propagation cleanup. ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations: - Canonicalize view/permute chain collection now uses deque + membership set instead of list pop(0) and remove(), avoiding quadratic bookkeeping. - FuseDuplicateUsersPass deduplicates pending producer revisits while preserving same revisit behavior after fusions. - PropagateViewCopyPermutePass still retraces after each moved transform for metadata safety, but defers horizontal/vertical cleanup until full scan finds no more direct propagation moves. Preserves fixed-point behavior while avoiding expensive cleanup after every single moved transform. Behavior-preserving speedup, all existing pass tests pass. Differential Revision: D114224762
bb32005 to
86beae5
Compare
Summary:
Speed up permute propagation cleanup.
ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations:
Behavior-preserving speedup, all existing pass tests pass.
Differential Revision: D114224762